Voice modeling methods for automatic speaker recognition

نویسنده

  • Thilo Stadelmann
چکیده

Building a voice model means to capture the characteristics of a speaker’s voice in a data structure. This data structure is then used by a computer for further processing, such as comparison with other voices. Voice modeling is a vital step in the process of automatic speaker recognition that itself is the foundation of several applied technologies: (a) biometric authentication, (b) speech recognition and (c) multimedia indexing. Several challenges arise in the context of automatic speaker recognition. First, there is the problem of data shortage, i.e., the unavailability of sufficiently long utterances for speaker recognition. It stems from the fact that the speech signal conveys different aspects of the sound in a single, one-dimensional time series: linguistic (what is said?), prosodic (how is it said?), individual (who said it?), locational (where is the speaker?) and emotional features of the speech sound itself (to name a few) are contained in the speech signal, as well as acoustic background information. To analyze a specific aspect of the sound regardless of the other aspects, analysis methods have to be applied to a specific time scale (length) of the signal in which this aspect stands out of the rest. For example, linguistic information (i.e., which phone or syllable has been uttered?) is found in very short time spans of only milliseconds of length. On the contrary, speakerspecific information emerges the better the longer the analyzed sound is. Long utterances, however, are not always available for analysis. Second, the speech signal is easily corrupted by background sound sources (noise, such as music or sound effects). Their characteristics tend to dominate a voice model, if present, such that model comparison might then be mainly due to background features instead of speaker characteristics. Current automatic speaker recognition works well under relatively constrained circumstances, such as studio recordings, or when prior knowledge on the number and identity of occurring speakers is available. Under more adverse conditions, such as in feature films or amateur material on the web, the achieved speaker recognition scores drop below a rate that is acceptable for an end user or for further processing. For example, the typical speaker turn duration of only one second and the sound effect background in cinematic movies render most current automatic analysis techniques useless. In this thesis, methods for voice modeling that are robust with respect to short utterances and background noise are presented. The aim is to facilitate movie

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A Critical Review on Automatic Speaker Recognition

Automatic Speaker Recognition (ASR) is use to recognizing persons from their voice. Since the voice of every human is not same because their vocal tract shapes, larynx sizes and other parts of a human voice production system. Automatic Speaker recognition is a procedure to automatically recognizing a speaker or who is speaking by the individual information counted in speech signal/waves. Automa...

متن کامل

Design of Matlab®-Based Automatic Speaker Recognition Systems

This paper presents design of an automatic speaker recognition system using Matlab® environment, which was part of a research project for NASA for undergraduate research experience. The project represents one of the many design and development activities that University of Maryland Eastern Shore offers as part of undergraduate research experience to undergraduate students in the area of Science...

متن کامل

Clustering Algorithm in Automatic Speaker Verification

We propose a new modeling approach in Automatic Speaker Verification A.S.V based on Gaussians Mixtures Models and Maximum a posteriori adaptation MAP. We propose clustering algorithm for intra and inter speaker’s variability in voice module and contribute for Universal Speaker Model design. We compare the traditional approach which uses one specific customer model with the second called Univers...

متن کامل

On Combining Classifiers for Password Secured Automatic Speaker Recognition System

Automatic Speaker recognition (ASR) is a pattern recognition problem that involves the process of automatically recognizing the speaker from their voices. Password protected speaker recognition system gives an extra security to the system where a person is not only identified by his natural voice biometric but also needs to remember a password (e.g. a combination lock number) that has to be spo...

متن کامل

Automatic Building of Synthetic Voices from Audio Books

Current state-of-the-art text-to-speech systems produce intelligible speech but lack the prosody of natural utterances. Building better models of prosody involves development of prosodically rich speech databases. However, development of such speech databases requires a large amount of effort and time. An alternative is to exploit story style monologues (long speech files) in audio books. These...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2010